PiPNN 3/6: add core graph construction - #1290
Conversation
There was a problem hiding this comment.
Pull request overview
This PR introduces the “core” PiPNN build pipeline in the diskann-pipnn crate, wiring together deterministic partitioning, leaf-local candidate construction, and final pruning into a public build_graph API with a validated build context.
Changes:
- Adds
PiPNNConfigvalidation and aPiPNNBuildContextthat binds PiPNN policy to DiskANN graph pruning policy and a caller-owned Rayon thread pool. - Implements the three main stages: partitioning (
partitioning.rs), leaf candidate construction (leaf_build.rs), and final pruning via shared Vamana robust prune (finalization.rs). - Adds comprehensive unit/integration tests and a Criterion benchmark for core scenarios; updates dependencies, lockfile, and mutation-test exclusions.
Reviewed changes
Copilot reviewed 13 out of 14 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| diskann-pipnn/src/lib.rs | Adds public PiPNN API (PiPNNConfig, PiPNNBuildContext, build_graph) and stage orchestration. |
| diskann-pipnn/src/partitioning.rs | Implements deterministic overlapping partition construction and leader assignment/scatter. |
| diskann-pipnn/src/partitioning/tests.rs | Adds unit tests covering partition determinism, invariants, error cases, and helpers. |
| diskann-pipnn/src/leaf_build.rs | Builds leaf-local symmetric k-NN candidates and accumulates global candidates safely in parallel. |
| diskann-pipnn/src/leaf_build/tests.rs | Adds unit tests for candidate correctness, invariants, type support, and error handling. |
| diskann-pipnn/src/finalization.rs | Orders/prunes candidate rows using shared robust_prune and validates candidate IDs/shape. |
| diskann-pipnn/src/finalization/tests.rs | Adds unit tests for pruning behavior and candidate validation failures. |
| diskann-pipnn/src/tests.rs | Tests effective_metric behavior for integer cosine-normalized handling. |
| diskann-pipnn/tests/config.rs | Integration tests for config validation and graph-policy compatibility checks. |
| diskann-pipnn/tests/build_graph.rs | Integration tests for end-to-end graph building, invariants, determinism, and type/metric support. |
| diskann-pipnn/benches/core.rs | Adds a Criterion benchmark for stage-focused core build scenarios. |
| diskann-pipnn/Cargo.toml | Updates crate dependencies/dev-dependencies and registers the new core benchmark target. |
| Cargo.lock | Records dependency graph changes for the updated diskann-pipnn crate dependencies. |
| .cargo/mutants.toml | Adds mutation-test exclusions for key PiPNN public boundary checks and partitioning invariants. |
Comments suppressed due to low confidence (1)
diskann-pipnn/src/partitioning.rs:604
size_of::<f32>()is used without being in scope (nouse std::mem::size_of;and not qualified), so this function won’t compile as written.
fn assignment_stripe_rows(leaders: usize) -> usize {
(ASSIGNMENT_CACHE_TARGET_BYTES / (leaders.max(1) * size_of::<f32>()))
.clamp(MIN_ASSIGNMENT_STRIPE_ROWS, MAX_ASSIGNMENT_STRIPE_ROWS)
}
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
5be0c9c to
50047c6
Compare
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 14 out of 15 changed files in this pull request and generated no new comments.
Comments suppressed due to low confidence (2)
diskann-pipnn/src/partitioning.rs:637
size_ofis used without being in scope (std::mem::size_of), which will not compile. Qualify the call or import it.
fn assignment_stripe_rows(leaders: usize) -> usize {
(ASSIGNMENT_CACHE_TARGET_BYTES / (leaders.max(1) * size_of::<f32>()))
.clamp(MIN_ASSIGNMENT_STRIPE_ROWS, MAX_ASSIGNMENT_STRIPE_ROWS)
}
diskann-pipnn/src/partitioning.rs:19
Normis imported but never used in this module, which will tripunused_importswarnings (and can become CI failures under-D warnings). Remove it from the import list.
use diskann::{utils::VectorRepr, ANNError, ANNResult};
use diskann_linalg::Transpose;
use diskann_utils::views::MatrixView;
use diskann_vector::{distance::Metric, norm::FastL2NormSquared, Norm};
use rand::{prelude::IndexedRandom, SeedableRng};
use rayon::prelude::*;
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 14 out of 15 changed files in this pull request and generated no new comments.
Comments suppressed due to low confidence (1)
diskann-pipnn/src/partitioning.rs:390
gather_rowsusesTypeId::of::<T>(), which implicitly requiresT: 'static. Making that bound explicit here avoids surprising/indirect trait-bound errors later and matches the publicbuild_graphboundary (which already requires'static).
fn gather_rows<T>(data: MatrixView<'_, T>, indices: &[u32], output: &mut [f32]) -> ANNResult<()>
where
T: VectorRepr,
1324668 to
857e200
Compare
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 14 out of 15 changed files in this pull request and generated 1 comment.
Comments suppressed due to low confidence (1)
diskann-pipnn/src/partitioning.rs:641
size_of::<f32>()is used without being imported or qualified, which will fail to compile. Qualify it withstd::mem::size_of(or add an explicit import).
let rows = ASSIGNMENT_CACHE_TARGET_BYTES / (leaders.max(1) * size_of::<f32>());
857e200 to
f642204
Compare
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 14 out of 15 changed files in this pull request and generated 1 comment.
Comments suppressed due to low confidence (4)
diskann-pipnn/src/leaf_build.rs:222
build_leafis executed from a Rayon parallel context (viabuild_leaf_candidates), so it should also explicitly requireT: Send + Syncto reflect the actual thread-safety requirement.
where
T: VectorRepr + 'static,
{
diskann-pipnn/src/leaf_build/tests.rs:154
assert_source_typeforwardsTinto the parallel leaf build path, so it should also includeSend + Syncbounds to match the production requirements.
fn assert_source_type<T>(data: &[T])
where
T: diskann::utils::VectorRepr + 'static,
{
diskann-pipnn/src/leaf_build.rs:193
build_leaf_candidatesuses Rayon parallel iteration overdata, soTmust beSend + Sync. Making this explicit in the signature avoids confusing trait-bound errors at call sites and documents the thread-safety requirement.
This issue also appears on line 220 of the same file.
where
T: VectorRepr + 'static,
{
diskann-pipnn/src/leaf_build/tests.rs:35
- This test helper calls
build_leaf_candidates, which (via Rayon) requiresT: Send + Sync. Add the bounds here so the test continues to compile once the production signature is tightened.
This issue also appears on line 151 of the same file.
where
T: diskann::utils::VectorRepr + 'static,
{
f642204 to
b1181d6
Compare
Pass the existing SortedNeighbors witness into internal prune so source-distance ordering is enforced by type rather than caller documentation.
Reject k outside 1..=3 at context construction so production never reaches an unsupported leaf-kernel width.
Partitioning is the only production source and emits sorted unique IDs. Reject unsorted input linearly and remove the HashSet fallback.
Rely on validated MatrixView shape and avoid cloning PiPNNConfig and its fanout vector.
Run partition orchestration under one architecture/metric specialization and reuse a runtime-sized tracker instead of imposing a fanout cap.
Keep architecture and metric concrete across the Rayon leaf pass and call the generic kernel directly.
| Transpose::Ordinary, | ||
| point_count, | ||
| leader_count, | ||
| dimensions, |
There was a problem hiding this comment.
Maybe I'm misunderstanding something, but it seems that this part of the computation scales linearly with the vector dimension. Given that PiPNN is mainly targeting faster index construction, do we expect this to become a bottleneck on high-dimensional datasets, or has that not been an issue in practice?
|
|
||
| diskann_linalg::sgemm_aat_lower( | ||
| point_ids.len(), | ||
| data.ncols(), |
There was a problem hiding this comment.
Similar to my understanding that the partitioning step is affected by dimensionality, the leaf-building stage here also seems to have computational complexity that scales roughly linearly with the dimension.
Purpose
This PR adds provider-independent PiPNN graph construction under
diskann::graph::pipnn.The core accepts a
MatrixView, graph policy, and Rayon pool. It returns one adjacency list for each data point.Main changes
PiPNNConfigdefines partition and leaf policy.build_graphselects one architecture and one metric marker.partitioningcreates overlapping bounded leaves.PartitionMetricto prepare point and leader norms.leaf_buildgathers vectors and computes lower-triangular Gram matrices.LeafMetricto prepare metric-specific norms.finalizationapplies shared RobustPrune only to overfull rows.The core does not load providers. It does not select start points. It does not serialize indexes.
Required invariants
u32.0 < c_min <= c_max.leaf_kis in the range 1 through 3.Review order
PiPNNConfig,PiPNNBuildContext, andbuild_graphinmod.rs.partitioning.rs.leaf_build.rs.finalization.rs.Validation
-Dwarnings.Stack
Stack 3/6. Depends on #1287. #1291 adds disk-index integration.